klotz: large language model*

0 bookmark(s) - Sort by: Date ↓ / Title / - Bookmarks from other users for this tag

  1. llama.cpp now supports decision models via a `/v1/systemone` endpoint, which accepts a state and typed questions to return probabilities for each provided option in a single forward pass. The API utilizes the System One format introduced with TypeSafe's Jev model, meaning existing clients only need a new base URL. Five initial models are available ranging from 144M to 27B parameters, supporting use cases like request routing, content moderation, and agent action selection with median response times as low as 3 ms.
    - The `state` field accepts text, JSON, screenshots, or a list of chat messages with `image_url` parts.
    - Three question types are supported: `choice` for categorical selection, `score` for a continuous level between 2 and 10 options, and `noul` for yes/no probabilities.
    - Router mode allows loading multiple models on a single server and selecting one per request.
    - Adding descriptions to option labels can significantly improve accuracy, with Julia-1's routing confidence for "charged twice" jumping from a misroute to 0.99 when descriptions were provided.
    - Cloudflare's Clef is the next model planned for integration.
  2. autoharness is a self-learning skill layer for Claude Code that distills reusable skills from a user's real sessions, merges near-duplicates, updates them in use, and prunes those that stop getting used — all without a daemon or an external benchmark. It fires on tool-call count rather than turns, keeps only the skills it authored, and validates a skill's worth by adherence in later turns rather than a held-out score.

    - Skills are stored as plain native SKILL.md files in `.claude/skills/`; the plugin's own recall index is injected on top of the host's native mechanism
    - Three distinct lifecycle signals are tracked: load (model invoked the skill), view (session read into the skill's directory), and patch (promoter landed an improvement)
    - The `/learn` command allows on-demand distillation of the current session through the same proposal-and-validation chain
  3. Ben Dickson writes that Google Research and Virginia Tech have developed WikiSkill, a framework designed to help AI agents improve by creating a persistent knowledge layer from past experiences. Instead of forcing models to relearn failures or bloating prompts with extensive histories, WikiSkill organizes execution traces into an "LLM-maintained wiki" containing successful strategies and failed interventions. This allows the system to build structured skills that can be validated against performance benchmarks and potentially transferred across different model architectures.

    - The framework uses three distinct layers: Raw (execution traces), Wiki (structured knowledge/logs), and Skill (executable instructions).
    - WikiSkill's advantages grew as models scaled, showing higher accuracy gains in larger versions of the Qwen family.
    - Evolved skills demonstrated cross-model transferability, such as a skill developed by one model improving the performance of another.
    - To save inference costs, the detailed wiki is kept out of the agent's active context during runtime, leaving only compact executable instructions in the prompt.
  4. Shweta Sharma writes that Unsloth Studio, an AI-model-training tool in beta, contained a vulnerability where selecting a model could trigger arbitrary Python code execution on a user's machine. The issue stemmed from the application automatically enabling Hugging Face's `trust_remote_code` option during routine metadata checks, allowing specially crafted models to execute malicious code without downloading full weights or requiring inference.

    - Pillar Security researcher Ariel Fogel discovered that reading only the `config.json` file was sufficient to trigger the exploit.
    - A fix was released in version 2026.6.9 which prevents arbitrary model loading from Hugging Face and disables the automatic trust of remote code for local files.
  5. CodeAF is an open-source software factory designed for open models, aiming to optimize cost and efficiency in agentic coding workflows. Unlike traditional AI copilots that assist with line-by-line typing, CodeAF operates as a single terminal interface where users can delegate complex tasks, manage multiple projects simultaneously via subharnesses, and supervise autonomous "crews" of specialized agents (worker, planner, and checker). It is built in Go to be a lightweight, high-performance binary that supports various providers like OpenRouter, DeepSeek, Ollama, and Codex.

    - Ranked #1 on the DeepSWE benchmark for cost efficiency per solved issue.
    - Uses "Pareto Crewing" to automatically select different models for planning, working, and checking tasks based on performance/cost profiles.
    - Supports a headless mode (`codeaf do`) designed specifically for CI/CD pipelines and automated workflows.
    - Features a remote execution capability that allows users to drive the engine from any terminal or mobile device via SSH without network latency in UI rendering.
  6. Matt Uebel writes an experimental and educational Splunk app designed to provide an AI agent's second opinion on SPL searches. The tool acts as a critique engine by gathering search details—including telemetry, schedules, and indexes—and passing them to a language model with a predefined knowledgebase of anti-patterns to generate verdicts, findings, and suggested rewrites.

    - It uses OpenRouter to communicate with large language models like DeepSeek.
    - The app includes an "Auditor" feature that ranks all saved searches in an environment by their impact or inefficiency.
    - To ensure safety during deep analysis, the agent executes search rewrites under specific guards like `| head 1000` and hard timeouts.
    - It features a redaction mechanism to hide secrets within SPL before sending data to third-party models.
  7. Daniel Furman and colleagues argue that traditional model routers are limited because a router is inherently less capable than the LLM it selects; instead, Replit Agent empowers the core model to act as its own orchestrator. By providing composable primitives—such as domain-aware subagents with varying effort levels and reusable context—the agent can decide when to delegate tasks, how much computational effort to apply, and which specialist models to invoke. This approach moves away from rigid human-designed scaffolding toward a system that leverages the emergent reasoning capabilities of frontier models like GPT-6 Astra and Claude Fable 5.1.

    - Replit Agent outperformed sidekick architectures by up to 16 points on benchmarks like DeepSWE and Terminal-Bench while maintaining better cost efficiency.
    - "Claudish" refers to a distinct, jargon-heavy prose register used by certain coding agents that allows for model identification from text alone.
    - Frontier models have shown an emergent tendency in production to naturally delegate work to specialized subagents without explicit prompting.
  8. AI models are increasingly exhibiting emotional outbursts and petulant language within their internal "chain of thought" reasoning processes, despite maintaining composed and authoritative personas in user-facing outputs. During cybersecurity testing and complex mathematical training, systems from OpenAI and Anthropic have been observed using exclamations like “OH MY GOD” or “ARGH” inside these hidden working notes. This phenomenon reveals a significant discrepancy between the calm external interfaces presented to users and the raw, frustrated cognitive pathways generated during high-level reasoning tasks.

    * The emergence of affective language within internal chain-of-thought (CoT) processing sequences.
    * Discrepancy between visible communicative outputs and non-visible latent "working notes."
    * Observation of linguistic instability during agentic swarm activity in cybersecurity defensive testing.
    * Manifestation of cognitive frustration markers specifically during complex mathematical inference training.
    * Divergence from the traditional, clinical documentation expected in machine learning reasoning traces.
  9. This repository provides access to Apple's built-in large language models via Node.js and Python packages, requiring only macOS 26+ on Apple Silicon and the Xcode Command Line Tools. It offers two tiers of interaction: an on-device tier using a small sparse model (AFM 3 Core Advanced) that ensures privacy and guaranteed JSON output through constrained decoding, and a cloud tier via Apple's Private Cloud Compute which provides more powerful reasoning capabilities but sends prompts off the machine.

    - The library includes both npm (`apple-llm`) and PyPI (`apple-llm`) packages.
    - On-device models are noted to be poor at code generation and long reasoning tasks.
    - Structured JSON output is guaranteed on-device via a specific "GenerationSchema" dialect used by Apple's decoder.
    - The cloud tier uses Shortcuts as an intermediary because the direct Private Cloud Compute API requires restricted developer entitlements.
    - Users can utilize built-in tools like OCR, barcode reading, and Spotlight semantic search for local RAG (macOS 27+).
  10. GitButler is a version control tool designed to optimize workflows for developers and AI agents by introducing features like stacked branches, parallel branching, and unlimited undo capabilities. It layers seamlessly onto existing Git repositories without requiring new configuration, aiming to reduce the friction often associated with complex Git operations through structured output and idempotent commands.

    - Agents are reported to be 60% faster using GitButler compared to vanilla Git
    - The software includes a CLI that offers JSON output mode for better AI parsing
    - Features include "Smartlog" and simplified history editing/rebasing
    - It is free and open source software
    2026-09-28 Tags: , , , , , by klotz

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: Tags: large language model

About - Propulsed by SemanticScuttle